Anthropic launches new model that imperceptibly watermarks texts made with their Language Models: ‘The load-bearing question is, how would anyone know if it was watermarked?’

Advertisement
  • New Watermark Rules

    *Claude
  • How would an “invisible watermark” in AI-generated text actually work?

    I saw that Anthropic is apparently planning to embed invisible watermarks into Claude- generated text, and this part caught my attention:
  • Does anyone here know how something like this actually works technically? If it's genuinely part of the text, I'm guessing it can't just be hidden Unicode characters, because those would be pretty easy to
  • detect and strip. So is it more like: choosing certain words/synonyms according to a statistical pattern?
  • slightly biasing token selection during generation? encoding a signal into the distribution of words or sentence structures?
  • something else entirely? And how could it survive editing? For example, if someone takes Claude's output, rewrites 20-30% of it,
  • changes sentence order, or runs it through another LLM, would the watermark still be detectable? I'm not asking about Al detectors in general. I'm specifically curious about the technical mechanism
  • All the AIs we use

    ΑΙ ChatGPT DeepSeek Claude Mistral Al Gemini Copilot 00 Poe
  • behind a watermark that's supposedly embedded in the text itself and survives copy/paste. Would love to hear from anyone who understands text watermarking or has read the relevant research.
  • FreakDeckard Here's how it works in simple steps: 1. The Al uses a secret rule when choosing words. Normally, an Al picks the next word based on probabilities. With a watermark, it slightly favors certain words that match a hidden pattern.
  • 2. This pattern is created using a secret key and the words that came before. The Al gives a small boost to words on a "good" list or with higher scores. 3. The change is so tiny that the text still sounds natural and makes sense. You can't tell by reading it.
  • 4. To check if text has a watermark, a detector uses the same secret rule to see if the text has more of the favored words than expected. It gives a score (like a z-score) to decide. 5. The watermark stays if you copy-paste the text or make small changes. But if you rewrite the text a lot (heavy paraphrasing), the watermark disappears.
  • 6. Different Al companies do it slightly differently: Google's Gemini uses a method called "tournament sampling." Claude (by Anthropic) also adds watermarks at the model level, according to their documents.
  • Human meets Robot

    Cheezburger Image 10656844032

Tags

Scroll Down For The Next Article